Papers with large-scale analysis

14 papers
Predicting Foreign Language Usage from English-Only Social Media Posts (N18-2)

Copied to clipboard

Challenge: Social media is known for its multi-cultural and multilingual interactions, a natural product of which is code-mixing.
Approach: They analyze 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English to build predictive models to infer non-English languages users speak exclusively from their tweets.
Outcome: The proposed models are based on a corpus of 6 million tweets produced by 27 thousand multilingual users speaking 12 other languages besides English . they show that content, style and syntax are the most predictive of non-English languages that users speak on Twitter.
WHoW: A Cross-domain Approach for Analysing Conversation Moderation (2025.naacl-long)

Copied to clipboard

Challenge: Using this framework, we annotated 5,657 sentences with human judges and 15,494 sentences with GPT-4o from two domains: TV debates and radio panel discussions.
Approach: They propose an evaluation framework for analyzing the facilitation strategies of moderators across different domains/scenarios by examining their motives (Why), dialogue acts (How) and target speaker (Who).
Outcome: The framework is generalisable across domains and reveals distinct modes of moderation: debate moderators emphasise coordination and facilitate interaction through questions and instructions, panel discussion moderator prioritize information provision and actively participate in discussions.
The structure of online social networks modulates the rate of lexical change (2021.naacl-main)

Copied to clipboard

Challenge: lexical change is a prevalent process, as new words are added, thrive, and decline in day-to-day usage.
Approach: They conduct a large-scale analysis of over 80k neologisms in 4420 online communities over a decade and found that the community’s network structure plays a significant role in lexical change.
Outcome: The results show that the community’s network structure plays a significant role in lexical change.
Modeling Framing in Immigration Discourse on Social Media (2021.naacl-main)

Copied to clipboard

Challenge: Using a dataset of immigration-related tweets, we examine how ordinary people on social media frame political issues.
Approach: They propose to use a dataset of immigration-related tweets labeled for multiple framing typologies from political communication theory to analyze framers.
Outcome: The proposed model enables comparisons between different types of frames on social media and a dataset of immigration-related tweets.
Large-Scale Hate Speech Detection with Cross-Domain Transfer (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for hate speech detection are limited due to the labor cost.
Approach: They construct large-scale tweet datasets for hate speech detection in English and a low-resource language, Turkish, consisting of human-labeled 100k tweets per each.
Outcome: The proposed datasets outperform conventional bag-of-words and neural models by at least 5% in English and 10% in Turkish for large-scale hate speech detection.
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance (2026.findings-eacl)

Copied to clipboard

Challenge: Large Language Models (LLMs) are often treated as defects of the model or its decoding strategy.
Approach: They construct a 22-dimension query feature vector covering clause complexity, lexical rarity, anaphora, negation, answerability, and intention grounding.
Outcome: The proposed model covers clause complexity, lexical rarity, anaphora, negation, answerability, and intention grounding, all known to affect human comprehension.
Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Temporal reasoning is a vital component of human communication and understanding, yet remains an underexplored area within the context of Large Language Models (LLMs).
Approach: They propose to use 3 prompting strategies to evaluate 8 different LLMs across 6 datasets and 2 Code Generation LMs to perform the analysis.
Outcome: The proposed models perform better on NLP tasks than the standard models on the same dataset.
Making “fetch” happen: The influence of social and linguistic context on nonstandard word growth and decline (D18-1)

Copied to clipboard

Challenge: In an online community, new words come and go, but language change is shaped and constrained by the grammatical system in which it takes part.
Approach: They analysed the frequency of non-standard words in reddit to determine their impact on language change.
Outcome: The results show that language change is shaped and constrained by the grammatical system in which it takes place.
OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants (2024.emnlp-main)

Copied to clipboard

Challenge: a large-scale analysis of millions of tweets on homelessness is challenging to understand at scale.
Approach: They propose a framing typology: Online Attitudes Towards Homelessness (OATH) They use large language models to analyze millions of tweets to find patterns in public attitudes .
Outcome: The proposed model speeds up annotations while incurring a 3 point performance reduction compared to existing classifiers .
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Understanding how jailbreaking works remains limited, hindering the development of effective defense strategies.
Approach: They propose a new mechanism that adaptively constrains activations within the safety boundary and propose 'Activation Boundary Defense' to enhance its effectiveness.
Outcome: The proposed defense achieves an average Defense Success Rate (DSR) of over 98% against various jailbreak attacks, with less than 2% impact on the model’s general capabilities.
MeasHalu: Mitigation of Scientific Measurement Hallucinations for Large Language Models with Enhanced Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) exhibit severe hallucinations, which undermine reliability of automated scientific document understanding systems.
Approach: They propose a framework for mitigating scientific measurement hallucinations through enhanced reasoning and targeted optimization.
Outcome: The proposed framework significantly reduces hallucination rates and improves overall accuracy on the MeasEval benchmark.
Are Language Models Consequentialist or Deontological Moral Reasoners? (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on the moral judgments in large language models rather than their underlying moral reasoning process.
Approach: They propose a taxonomy of moral rationales to classify reasoning traces according to consequentialism and deontology . they use trolley problems to analyze moral reasoning tracing in large language models .
Outcome: The proposed taxonomy of moral rationales sheds light on consequentialism and deontology . it systematically classifies reasoning traces according to two main ethical theories .
A Scalable Entity-Based Framework for Auditing Bias in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing approaches to bias evaluation in large language models trade ecological validity for statistical control, or use artificial prompts that lack scale and rigor.
Approach: They propose a framework that uses named entities as probes to measure bias in large language models.
Outcome: The proposed framework reproduces bias patterns observed in natural text, enabling large-scale analysis.
Multilinguality Does not Make Sense: Investigating Factors Behind Zero-Shot Cross-Lingual Transfer in Sense-Aware Tasks (2025.emnlp-main)

Copied to clipboard

Challenge: Cross-lingual transfer allows models to perform tasks in languages unseen during training and is often assumed to benefit from increased multilinguality.
Approach: They challenge this assumption by analyzing polysemy disambiguation and lexical semantic change in 28 languages and using confounding factors to account for perceived advantages.
Outcome: The proposed models and benchmarks are compared across 28 languages and show that multilingual training is neither necessary nor beneficial for effective transfer.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations